Goto

Collaborating Authors

 efficient training and inference


The 2 Types of Hardware Architectures for Efficient Training and Inference of Deep Neural Networks

#artificialintelligence

Due to the popularity of deep neural networks, many recent hardware platforms have special features that target deep neural network processing. The Intel Knights Mill CPU will feature special vector instructions for deep learning. The Nvidia PASCAL GP100 GPU features 16-b floating-point (FP16) arithmetic support to perform two FP16 operations on a single-precision core for faster deep learning computation. Systems have also been built specifically for DNN processing, such as the Nvidia DGX-1 and Facebook's Big Basin custom DNN server. DNN inference has also been demonstrated on various embedded System-on-Chips (SoCs) such as Nvidia Tegra and Samsung Exynos, as well as on field-programmable gate arrays (FPGAs).